Your medical records are protected by law. The data about your health mostly is not.
That single distinction explains more about healthcare privacy than any discussion of encryption or consent forms. The notes your doctor writes, the results of your blood test, the images from your scan — these sit inside a regulated system with legal obligations attached. Meanwhile the fitness app tracking your sleep, the search you made about a symptom, the pharmacy loyalty card recording what you buy, and the period tracker on your phone generally sit outside it entirely.
Understanding where that line falls is the most practically useful thing anyone can know about AI healthcare privacy, because it determines which protections you actually have.
Covered and Uncovered
Broadly, and with variation between countries:
Usually covered by health privacy law:
- Records held by hospitals, clinics, and doctors.
- Laboratory and imaging results from clinical providers.
- Prescription records held by pharmacies in their clinical role.
- Insurance claims data, in insurance-based systems.
Usually not covered:
- Consumer fitness trackers and wellness apps.
- Symptom checkers and general health websites.
- Direct-to-consumer genetic testing, in many jurisdictions.
- Search history and social media activity revealing health concerns.
- Retail purchase data.
- General-purpose chatbots used to ask health questions.
The second list frequently contains more revealing information than the first. A person's search history often shows what they were worried about months before they mentioned it to a doctor.
Europe's general data protection framework is broader than the American sectoral approach and treats health data as a special category regardless of who holds it. Other jurisdictions vary widely. Where you live materially changes what protections apply, which is worth knowing before assuming any.
De-identification Is Weaker Than It Sounds
When health data is shared for research or commercial purposes, it is typically described as de-identified or anonymised. Names removed, identifiers stripped.
Researchers have repeatedly demonstrated that this is less protective than it appears. A combination of ordinary attributes — postcode, date of birth, sex, dates of admission — can narrow a record to a single individual with surprising frequency. Genomic data is inherently identifying and cannot be meaningfully anonymised at all. Continuous data such as movement or heart rhythm patterns can function as a signature.
This does not mean data sharing is illegitimate. Research requires data, and medical progress depends on it. It means that "anonymised" should be read as "harder to identify," not as "impossible to identify" — and that the protections which matter are legal and organisational rather than purely technical.
Where Your Data Actually Goes
Most patients have no clear picture of this, and the honest answer is that several flows exist simultaneously:
- Your direct care. Uncontroversial and expected.
- Internal quality improvement, generally permitted without specific consent.
- Research, under governance arrangements that vary by country and institution.
- Commercial partnerships between health systems and technology companies, which have generated significant controversy in several countries.
- Model training, where clinical data is used to develop systems that may later be sold commercially.
- Secondary data markets, where de-identified health data is traded, legally, in several jurisdictions.
The fifth is the newest and least understood by patients. Data collected during care, contributed without payment, may be used to build a product subsequently sold back to the health system that supplied the data. Whether that is acceptable is a legitimate question, and it is being asked in several countries.
The Consent Problem
Consent frameworks designed for a single research study translate poorly to continuous data use at scale.
The difficulties are structural rather than administrative:
- Broad consent asks people to agree to unspecified future uses, which is difficult to call informed.
- Consent fatigue is real. Nobody reads the twelfth form.
- Refusal often is not practical. Declining data sharing may mean declining care.
- Withdrawal is frequently impossible once data has been incorporated into a trained model.
- Group harms escape individual consent. Data about people like you can affect you even if you personally consented to nothing.
That final point is underappreciated. Predictions made about a group — a postcode, an ethnic group, an age band — affect individuals within it who never contributed anything.
Bias Is an Ethical Issue, Not a Technical One
A well-documented 2019 study in Science found that an algorithm used widely across the American health system to identify patients needing additional care systematically under-referred Black patients relative to their actual illness. The mechanism was that the model used healthcare spending as a proxy for healthcare need, and historically less had been spent on those patients at equivalent levels of illness.
The important feature of that case is that nobody intended it, race was not an input, and the algorithm was doing exactly what it was built to do. The unfairness came from the choice of proxy, made early and without much scrutiny.
What patients and advocates can reasonably ask:
- Which populations was this system validated on?
- Has performance been examined separately across demographic groups?
- What outcome is it actually predicting, and is that the outcome we care about?
- Who is monitoring for differential performance after deployment, and how often?
Explainability and the Right to Know
If a system influences your care, can anyone explain how?
For some models the answer is genuinely no in any detailed sense, which creates a real tension. A clinician can explain their reasoning; a complex model often cannot be reduced to one.
European regulation has moved toward classifying medical applications as high-risk, with obligations around transparency, human oversight, and documentation. Other jurisdictions are developing their own approaches at varying speeds.
The practical protection for patients is less about technical interpretability and more about process: a clinician who has reviewed the output, understands its limitations, and can explain the decision they made — including where they disagreed with the system.
Who Is Responsible When It Is Wrong
This remains unsettled, and patients should know that it is unsettled.
The candidates for responsibility are the clinician who acted on the output, the institution that deployed the system, the developer who built it, and the regulator who authorised it. Different legal systems are landing in different places, and case law is thin.
The current practical position in most settings is that the clinician retains responsibility, which is one reason assistive rather than autonomous deployment remains the norm. A clinician who is accountable for an output they cannot inspect is in an uncomfortable position, and that discomfort is doing useful work in keeping humans in the loop.
The Downstream Uses Patients Worry About Most
When people express unease about health data, they are rarely worried about researchers. They are worried about insurers, employers, and immigration authorities.
Those concerns are not unfounded, and the protections vary enormously:
- Insurance. Several countries restrict the use of genetic information in underwriting, sometimes through legislation and sometimes through voluntary industry agreements. Coverage of other health data is patchier, and life and income protection insurance is often treated differently from health insurance.
- Employment. Workplace wellness programmes that collect health data occupy an uncomfortable position, particularly where participation affects benefits or where the employer receives anything beyond aggregate figures.
- Immigration and visas. Health disclosure requirements exist in many countries, and data held elsewhere may be requested.
- Law enforcement. Access to health records varies by jurisdiction and by circumstance, and consumer app data is generally easier to obtain than clinical records.
- Divorce and custody proceedings, where health information is sometimes sought as evidence.
The practical implication is not to avoid healthcare. It is to be deliberate about which data goes into consumer applications with weak protections, particularly for conditions that carry social or legal consequences — mental health, reproductive health, substance use, and infectious disease among them.
Questions You Can Actually Ask
Patients often assume these questions are unwelcome. In well-run institutions they are not.
- Is any automated system involved in my care? Reasonable, and increasingly relevant.
- Is this consultation being recorded or transcribed? You are generally entitled to know, and often to decline.
- What happens to my data beyond my treatment?
- Can I opt out of research use without affecting my care?
- Who else has access to my record, and is access logged?
- How do I obtain a copy of my own records?
- How do I request correction or deletion, and what are the limits on that?
You may not always get complete answers. Asking still matters, because institutions respond to what patients ask about.
What Good Practice Looks Like
For anyone evaluating an institution or a service:
- Clear, readable explanation of data use — not buried in a fourteen-page document.
- Genuine opt-outs that do not penalise the patient.
- Access logging with audit, so inappropriate access is detectable.
- Defined retention periods, enforced.
- Independent ethics review of new deployments, not only of research.
- Monitoring of performance across demographic groups after go-live, not just before.
- A named person accountable, whom a patient can actually reach.
The Balance Worth Striking
There is a genuine tension here that privacy advocates sometimes understate. Data enables research, and research produces treatments. A world with maximal privacy protection and no data sharing would be a world with slower medical progress, and that cost falls on future patients who cannot advocate for themselves.
The reasonable position is not maximal restriction. It is proportionality, transparency, and accountability: data used for purposes patients would recognise as legitimate, with governance that is visible, and with someone answerable when it goes wrong.
What makes AI healthcare privacy difficult is that the current arrangement often fails on the middle requirement. Patients are not told, in terms they can understand, what is happening with information about their bodies.
That is fixable, and it does not require restricting research. It requires institutions to explain themselves — which is a lower bar than the current situation suggests.

0 Comments